[CI/Build] Add self-hosted MI350X ROCm external PD tier - #3338
Conversation
f59afd9 to
c59e3d6
Compare
c59e3d6 to
6e26005
Compare
|
Hi @andyluo7 Is this PR ready to merge? |
|
@staryxchen waiting for Wenjie's confirm. |
Hi @andyluo7, thanks for the work on this. I'd like to see this PR merged — the MI350X external-PD tier fills a real gap in our CI coverage for the ROCm path. Separately, since you already added the ROCm release pipeline in #3184: could you help push the mooncake-transfer-engine-rocm wheel to PyPI as well? CUDA/EFA/MUSA/NPU already have their variants published there, but the pipeline hasn't been triggered yet. |
Aionw
left a comment
There was a problem hiding this comment.
Hello, Stary. I've been validating this PR on a separate branch, where most of the issues have already been fixed. I'll port those changes back to this PR once the vLLM-related issue is resolved.
# Conflicts: # .github/workflows/ci.yml
# Conflicts: # .github/workflows/ci.yml
* [TransferEngine] Keep RDMA owner CQ polled during active handshake (#3814)
* fix: drain RDMA CQ during active handshake
* Fix comments
* polling for initial step
* [TENT] Fix DeviceSelector inflight quota accounting (#3511)
* [TENT] Fix DeviceSelector inflight quota accounting
* Update mooncake-transfer-engine/tent/tests/quota_accounting_test.cpp
Co-authored-by: Stary <staryxchen@tencent.com>
---------
Co-authored-by: jipengtian <jipengtian@xiaohongshu.com>
Co-authored-by: Stary <staryxchen@tencent.com>
* [PG] Fix CUDA enqueue-stream serialization (#3798)
* [PG] Add support for torch==2.14.0 (#3842)
* [Bugfix][TE] Use the only ADXL endpoint when dest has one engine (#3835)
Embedded/standalone TEs stamp buffer device_id with the logical NPU id
but publish a single endpoint at index 0, which spammed fallback
warnings. Multi-endpoint dest routing is unchanged.
Co-authored-by: youyx <youxiao@huawei.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* [Bugfix][TENT] Keep RailMonitor state across COW topology refreshes (#3832)
Segment metadata is copy-on-write, so buffer/CUDA-graph register
publishes a new Topology* even when the rail map is unchanged. Treat
same NIC/memory layout as a pin refresh instead of loadDefault(), and
log the config banner only on first load so PD/e2e is not flooded.
Signed-off-by: staryxchen <staryxchen@tencent.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* [TENT]: Deduplicate TentMetrics singleton via shared tent_metrics library (#3797)
* fix-tent-metrics-duplicate-log
* delete test
* link Rust TENT
* Fix tent_metrics RPATH and drop redundant Rust link paths
* Restore tent_metrics Rust link search path
---------
Co-authored-by: ruanzhao <ruanzhao@kingsoft.com>
* [TENT] Route HP TCP lanes across configured rails (#3834)
* feat(tent): route HP TCP lanes across rails
* fix(tent): validate HP TCP rail listener
* [Store] Preserve replica IDs during standby restore (#3811)
Co-authored-by: Yuchen Kou <kouyuchen@approaching.ai>
* [Store] Enable fenced batch OpLog writer (#3810)
* [Store] Enable fenced batch OpLog writer
* [Store] Preserve direct OpLog startup while fencing HA serving
---------
Co-authored-by: Yuchen Kou <kouyuchen@approaching.ai>
* [Store] Structure ReplicaSelectionConfig environment settings (#3840)
Signed-off-by: Schatten <czhengt@qq.com>
* [Store] Align OffsetAllocatorBackendConfig with config layout (#3823)
Co-authored-by: Aoi <aione.moe.dev@gmail.com>
Signed-off-by: Schatten <czhengt@qq.com>
* fix(te): free wr_depth_list_ of zero-QP endpoints in ~RdmaEndPoint (#3846)
construct() allocates wr_depth_list_ even when num_qp_list is 0, but the
destructor's abnormal-shutdown fallback only ran deconstructLocked() for
a non-empty qp_list_, leaking the array for constructed zero-QP
endpoints. Key the fallback on wr_depth_list_ too; deconstructLocked()
already handles the empty-qp_list_ case and nulls the pointer.
The nightly-coverage ASan build has failed rdma_endpoint_state_test at
exit on exactly this since the zero-QP lifecycle tests landed.
Fixes #3839.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* test: migrate placement, replication, and task scenarios to the DSL (#3702)
Move the client-visible placement/allocation, copy/move/replica-clear,
task RPC, graceful-unmount, and BatchQueryIp coverage from
master_service_test.cpp into four scenario suites on the MasterScenario
DSL, consuming the vocabulary from #3510 without extending it. Lease
and grace-timer sleeps become deterministic ExpireAt and Eventually.
Drain jobs, task payload wire format, private worker/refcount/metric
probes, and eviction-pressure tests stay as focused direct tests; a
trimmed MoveEndReleasesSourceRefcountWhenTargetGone direct test keeps
the private refcount half of the legacy MoveEnd case.
Implements #3573. Stacked on #3656.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* [Store] Complete KV event publication (#3818)
* [Store] Complete KV event publication
Finish the master-side KV cache event publisher so external cache-aware
indexers can consume a correct availability stream.
* Carry replica availability per medium instead of collapsing an object's
metadata into a single medium string, so a tier is never dropped from the
stream and a partial eviction retracts only the tier that went away.
* Announce evictions that leave the object alive, and pick a single publisher
per removal so a retraction is not emitted twice.
* Add a tenant-wide clear event, tied to the real outcome of a removal.
* Stamp the configured model context on every event instead of nil.
* Add publisher and master-service test coverage, and document the wire format
and the subscriber contract.
Co-Authored-By: lolo-pop <lolo-pop.rong@gmail.com>
* [Bugfix]fix removeall potential concurrency conflict
---------
Co-authored-by: j00628475 <jiangzexu144@huawei.com>
Co-authored-by: lolo-pop <lolo-pop.rong@gmail.com>
* [Store] Introduce placement index and replica allocator (#3704)
* [Store] Introduce stateful region resource drivers
* [Store] Introduce placement index and replica allocator
* [Store] Refine placement index and policy abstractions
* [Store] Preserve CXL-only preferred allocation
* [Store] Make preferred target filtering explicit
* [Store] Partition placement groups by target kind
* [Store] Make placement group kind explicit
* [Store] Self-heal dangling LOCAL_DISK replicas in Client::Put (#3711)
* [Store] Self-heal dangling LOCAL_DISK replicas in Client::Put
RemoveAll wipes client SSD files unconditionally while master keeps
lease-protected metadata, leaving a completed LOCAL_DISK entry whose
backing file is gone (issue #3709). Reads then fail, and every later Put
for the key hit OBJECT_ALREADY_EXISTS, which Client::Put converted into
a silent success, so the engine believed data was stored when nothing
was ever written.
When PutStart reports OBJECT_ALREADY_EXISTS and every completed replica
is a LOCAL_DISK replica owned by this client, verify the file with the
local storage backend; when it is gone, evict the dangling replica
through the existing EvictDiskReplica RPC and retry PutStart once. If
any completed replica lives elsewhere (memory, remote disk, DFS) the key
is still readable, so the old idempotent path is kept unchanged.
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
* [P2P-Store] Unregister buffers on GetReplica failure paths (#3783)
* [P2P-Store] Unregister buffers on GetReplica failure paths
doGetReplica left earlier registrations behind when a later Add or a
transfer failed, and GetReplica leaked all of them when metadata
recheck or the final update failed.
Signed-off-by: chlins <chlins.zhang@gmail.com>
* Address review: wait for in-flight transfers before rollback, undo per-round refcounts
Signed-off-by: chlins <chlins.zhang@gmail.com>
---------
Signed-off-by: chlins <chlins.zhang@gmail.com>
* [Bugfix] Fail SSD eviction tests on read errors (#3738)
* test: fail SSD eviction checks on read errors
Keep cleanup and diagnostics, then fail single and batched readback tests when data is empty or corrupted.
Assisted-by: Claude
Signed-off-by: Miguel Garcia <miguelgarciaroman8@gmail.com>
* test: isolate SSD eviction keys across cases
Use a distinct key prefix per test so asynchronous cleanup cannot leak stale SSD replicas into the next case.
Assisted-by: Claude
Signed-off-by: Miguel Garcia <miguelgarciaroman8@gmail.com>
* test: fail concurrent SSD reads on errors
Assert the aggregated concurrent read failure count after cleanup so empty or corrupted reads cannot leave the stress test green.
Assisted-by: Claude
Signed-off-by: Miguel Garcia <miguelgarciaroman8@gmail.com>
---------
Signed-off-by: Miguel Garcia <miguelgarciaroman8@gmail.com>
* [P2P-Store] Release batch IDs on all performTransfer exit paths (#3770)
* [P2P-Store] Release batch IDs on all performTransfer exit paths
performTransfer leaked the allocated batch ID on five paths: the
location==nil break, openSegment/submitTransfer/getTransferStatus
errors, and context cancellation.
Signed-off-by: chlins <chlins.zhang@gmail.com>
* Drain in-flight task on cancellation before freeing batch ID
freeBatchID returns BatchBusy while a task is pending, so the
cancellation path must reach a terminal state first. Extract a
batchTransport interface and add cancellation regression tests.
Signed-off-by: chlins <chlins.zhang@gmail.com>
---------
Signed-off-by: chlins <chlins.zhang@gmail.com>
* [Build] conductor: build skeleton and common layer (#3825)
* [Build] conductor: build skeleton and common layer
First slice of the Mooncake Conductor upstreaming series. Adds the
component skeleton so the later slices are pure additions:
- mooncake-conductor/CMakeLists.txt building the conductor_cpp_core
static library, with a standalone-build guard so the component also
configures on its own (outside the top-level GLOBAL_CONFIG build).
- common layer: ServiceConfig / HashProfileConfig data model
(common/types.h) and the environment-variable and log-level helpers
(common/utils.{h,cpp}).
- tests/CMakeLists.txt with the conductor_test gtest target, covering
the env helpers and the JSON uint64 round-trip contract that the
service HTTP responses rely on.
- WITH_CONDUCTOR option (default OFF) wired into the top-level build,
enabled in the CI unit-test configure so conductor_test is picked up
by the existing unified CTest run, and documented in build.md.
No conductor binary is produced yet; src/main.cpp arrives with the KV
event manager slice.
Co-Authored-By: Misak2333 <167268798+Misak2333@users.noreply.github.com>
* [Build] conductor: use the top-level googletest targets when present
* [Build] conductor: nest common namespace under mooncake
Address review feedback on #3825: use mooncake::conductor::common instead
of a top-level conductor namespace, matching the convention used across the
tree (mooncake::ha, mooncake::tent::ub, mooncake::offset_allocator).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* [Build] conductor: reuse mooncake-common parsing helpers
---------
Co-authored-by: rongchenghao <r00875518@china.huawei.com>
Co-authored-by: Misak2333 <167268798+Misak2333@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* [TENT] Add native UB endpoint bootstrap and READ/WRITE data path (#3796)
* [TENT] Add native UB endpoint bootstrap and READ/WRITE data path
* [TENT] Prefer UB bonding devices when device_filter is empty
Avoid initializing underlying udmac* contexts alongside bonding_dev_0
by default, which can bypass UBAGG and cause cross-node LOC_ACCESS_ERR.
Co-authored-by: Cursor <cursoragent@cursor.com>
* [TENT] Port UB bonding runtime patches into urma_adapter
Bring over the local working-tree fixes needed for bonding_dev_0:
SET_BONDING_MODE, CTP priority/multi_path jetty config, and token_id
segment registration. Also wire topology discover_ub from ub/enable.
Co-authored-by: Cursor <cursoragent@cursor.com>
* [TENT] Unify UB bonding detection and harden segment teardown
Reuse device_selection helpers in urma_adapter and install path so heuristics cannot drift, and always free token_id when unregister fails.
Co-authored-by: Cursor <cursoragent@cursor.com>
* [TENT] Fix clang-format for UB transport files
Apply clang-format-20 to changed lines in ub_transport.cpp and
urma_adapter.cpp so CI code format check passes.
Co-authored-by: Cursor <cursoragent@cursor.com>
* [CI/Build] Ignore gflags exit code in UB tebench smoke test
tebench --help exits 1 under gflags; tolerate that so the UB benchmark
CLI smoke step checks help text instead of failing early under set -e.
Co-authored-by: Cursor <cursoragent@cursor.com>
* [CI/Build] Restore UB tebench smoke step after main sync
Re-apply the ub-mock CLI help smoke test that was dropped while
resolving the merge onto the latest LinQuick main.
Co-authored-by: Cursor <cursoragent@cursor.com>
* [TENT][UB] Address PR #3796 review: reclaim provider-lost inflight token, decouple Topology from UB enable
- scanTimeouts(): after a successful quiesce fence, explicitly remove and
release a still-present timed-out token before scheduling its retry. A
provider-lost completion previously left the token in inflight_, so
releaseInflight() never ran and endpoint outstanding/quota stayed charged
until stop(), which could exhaust capacity and stall later transfers.
- Topology::discover(): drop the transport-policy discover_ub flag and keep
discovery transport-agnostic (always attempt UB when the platform supports
it); UbTransport activation remains gated by transports/ub/enable in the
transport loader.
- Add regression test ProviderLostCompletionIsReclaimedAfterSuccessfulFence
with a fake-adapter switch that makes quiesce succeed without returning the
token, asserting inflight/quota/outstanding accounting return to zero and
subsequent transfers still make progress.
Co-authored-by: Cursor <cursoragent@cursor.com>
* [TENT][UB] Fix clang-format in ub_native_data_path_test
Co-authored-by: Cursor <cursoragent@cursor.com>
* [TENT][UB] Cap slices per task and harden endpoint cleanup against incomplete teardown
- Add transports/ub/max_slices_per_task (default 64k) and reject requests
whose slice count exceeds it in submitTransferTasks, bounding per-task
memory/scheduling footprint.
- EndpointStore: treat retire() as incomplete when the endpoint does not
reach kDestroyed, quarantine it instead of replacing it, and retry
getOrCreate() so a same-key endpoint is not reused while cleanup is still
in progress.
- Tests: fake adapter gains reset-failure injection; new
EndpointStoreDoesNotReplaceEndpointWhenCleanupCannotComplete plus slice-cap
rejection coverage.
Co-authored-by: Cursor <cursoragent@cursor.com>
* [TENT][UB] Fix clang-format in tent_backend
* [TENT][UB] Remove ineffective tebench --help smoke step from CI
The step only grepped static gflags help strings and never exercised
the TENT/UB path; ub-mock build and CTest already cover real coverage.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Le1zyCatt <2490162471@qq.com>
Co-authored-by: Connor-Matthew <60215777+Connor-Matthew@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* Bump google.golang.org/grpc in /mooncake-p2p-store/src/p2pstore (#3827)
Bumps [google.golang.org/grpc](https://github.com/grpc/grpc-go) from 1.79.3 to 1.83.1.
- [Release notes](https://github.com/grpc/grpc-go/releases)
- [Commits](https://github.com/grpc/grpc-go/compare/v1.79.3...v1.83.1)
---
updated-dependencies:
- dependency-name: google.golang.org/grpc
dependency-version: 1.83.1
dependency-type: indirect
...
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
* [Bugfix][TENT] Return an error for unregistered application RPC IDs (#3862)
* [TENT] Slice large HP TCP reads across rails (#3861)
* [TENT] Slice large HP TCP reads across rails
* [TENT] Simplify HP TCP read slicing state
* [TENT] Cancel sibling HP TCP slices on first error
* [TENT] Keep sibling cancellation fix minimal
* [TENT] Test sliced failure cancellation and retry
* [TENT] Make sliced cancellation test deterministic
* [CI] Separate concurrent SSD write and read phases (#3870)
* [Reshard] Add Store-backed model-weight snapshots (#3772)
This PR adds Store-backed model-weight snapshots for heterogeneous resharding.
WeightStore persists a complete, immutable StoredWeightManifest only after
all planned payload fragments have been uploaded and verified. The manifest is
address-free and retains the logical tensor/fragment layout required to plan a
later restore into a different target topology. Runtime GPU addresses,
allocation bounds, leases, and allocation guards stay in ephemeral runtime
bindings.
The Store path provides:
generation-consistent source replica selection and upload planning;
grouped payload and metadata records through Mooncake Store;
commit/abort decisions, payload completeness checks, and recovery-safe
registration cleanup;
retry-safe manifest publication after a commit decision;
range-based restore with get_into_ranges into validated target bindings;
an explicit WeightStoreWriter lifecycle:
begin_weight_snapshot(...) -> write_tensor(...) -> commit().
* [TransferEngine] Clear pending notify when freeing batch (#3891)
* fix(te): clear pending notify when freeing batch
* fix(te): narrow pending notify lock scope
---------
Co-authored-by: guorongjie <guorongjie@bytedance.com>
* [Bugfix][TENT] Predict deadline MLU from a metered wire rate and the bytes ahead (#3780)
* [Bugfix][TENT] Return every slice's charge and lane count to the counter that holds it
* [Bugfix][TENT] Meter the NIC's transmit rate and predict the admission drop from it and the bytes ahead
* [TENT] Score each arbitration slot against the bytes actually posted ahead of it
---------
Co-authored-by: maxlisongsong <maxlisongsong@didiglobal.com>
* [Store] Support Dummy external buffer registration (#3721)
* test: migrate storage-tier RPC scenarios to the DSL (#3847)
Move the client-visible DFS, root_fs_dir dual-replica, and local-disk
offload coverage from master_service_test.cpp, master_service_ssd_test.cpp
and offload_on_evict_test.cpp into three scenario suites, adding the
local-disk vocabulary proposed on the issue (MountLocalDisk,
UnmountLocalDisk, ReportSsdCapacity, OffloadHeartbeat, CompleteOffload,
EvictDiskReplica) as thin wrappers over exported RPCs, each covered by
failure-reporting contract tests.
Metric-singleton probes, the PushOffloadingQueue private hook,
eviction-pressure and concurrency tests, the perf comparison, and the
USE_NOF cases stay as focused direct tests.
Closes #3574.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* [CI/Build] Add Kylin support to dependencies installer (#3906)
Co-authored-by: LiChangXin-KylinSoft <vovalcx@gmail.com>
* [CI/Build] Upload pre-release wheels to GitHub Releases (#3908)
* [Store] Expose explicit soft pin action and TTL in C ABI and Go/Rust bindings (#3881)
* [Store] Expose explicit soft pin action and TTL in C ABI and Go/Rust bindings
Complete the client-API migration started by the fixed soft pin
lifecycle: the C ABI boolean with_soft_pin becomes an explicit
mooncake_soft_pin_action plus an optional request-level TTL, and the
Go and Rust wrappers expose SoftPinAction/soft_pin_ttl_ms accordingly.
This lets non-C++ clients enable a soft pin with a custom TTL and
disable an active one instead of only toggling the master default.
Unknown actions and invalid action/TTL combinations are rejected
(wrapper-level in Go/Rust, master-side for raw C callers).
* test(store/go): allow overriding integration test endpoints via env
Respect MC_METADATA_SERVER and MOONCAKE_MASTER so the Go integration
tests can run against a master on non-default ports; defaults are
unchanged.
* style(store/go): rename SoftPinTTLMS to SoftPinTTLMs
Match the existing Go millisecond-suffix convention (timeoutMs in
mooncake-common).
* fix(store): reject out-of-range soft pin actions at the C ABI boundary
The C soft_pin_action field is a plain int while the C++ enum is
uint8_t-backed, so a direct static_cast wrapped large values (e.g.
257) into valid actions instead of letting the master reject them.
Validate against the enumerator range and fail with -1 like other
argument errors.
* [TENT] Disable Nagle on HP TCP client sockets (#3883)
* [TENT] Pin dispatching_bytes_ to the eligible bytes in Dispatching (#3910)
Co-authored-by: maxlisongsong <maxlisongsong@didiglobal.com>
* [TENT] Drain RDMA CQ and dest GPU before tearing down MRs (#3899)
* [TENT] Drain RDMA CQ and dest GPU before tearing down MRs
Quiesce stops new submits, waits for inflight to hit zero, then
cudaDeviceSynchronize on topology CUDA devices so GPUDirect writes are
visible before deconstruct unregisters memory.
Signed-off-by: staryxchen <staryxchen@tencent.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* [TENT] Move dest-GPU sync onto Platform during RDMA quiesce
Keep vendor CUDA/HIP calls out of RDMA transport. synchronizeDevices()
lives on Platform; CUDA and ROCM implement it, other platforms no-op.
Signed-off-by: staryxchen <staryxchen@tencent.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Signed-off-by: staryxchen <staryxchen@tencent.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* [Store] Add bounded standby promotion handoff (#3841)
Co-authored-by: Yuchen Kou <kouyuchen@approaching.ai>
* [Store] Structure TransferSubmitterConfig environment settings (#3850)
Signed-off-by: Schatten <czhengt@qq.com>
* [Store] Structure LocalFileSnapshotConfig environment settings (#3872)
Co-authored-by: Aoi <aione.moe.dev@gmail.com>
Signed-off-by: Schatten <czhengt@qq.com>
* [Store] Structure RpcTimeoutConfig environment settings (#3871)
Co-authored-by: Aoi <aione.moe.dev@gmail.com>
Signed-off-by: Schatten <czhengt@qq.com>
* [Store] Structure FilePerKeyConfig environment settings (#3853)
Signed-off-by: Schatten <czhengt@qq.com>
* [Feature] Add per-file storage mode to storage_benchmark_v1 (#3454)
* [Store] Stabilize HA under OpLog pressure (#3860)
Signed-off-by: Zhen Zhang <zhenzhang@roblox.com>
* [CI/Build] Run current weight snapshot API tests in nightly (#3930)
Signed-off-by: Miguel Garcia <miguelgarciaroman8@gmail.com>
* [Bugfix][TransferEngine] Prevent peer failures from latching RDMA contexts inactive (#3559)
* [Bugfix][TransferEngine] Harden the RDMA context breaker
Incident (2026-07-15): one mooncake store replica restarted during an mlx
fault. Five prefill pods accumulated 32 consecutive all-rails-failed submit
batches, every batch targeting that single store, and tripped the
context-level circuit breaker from #2155. The local RdmaContext became
inactive even though its port was still up. Since inactive contexts stop
posting work, CQ success could not clear the breaker, and PORT_ACTIVE never
arrived to start physical recovery. Every later PUT failed synchronously with
AddressNotRegistered until the processes restarted.
Make the context breaker distinguish peer failures from local RNIC failure:
1. Require an all-rails-failed submit streak to span at least
MC_CONTEXT_FAILURE_MIN_PEERS distinct peer servers before it may trip
(default 2; 1 restores single-peer submit tripping). Direct local WC
failures remain strong local evidence and preserve their existing
single-peer threshold behavior from #3387.
2. Auto-reactivate a breaker-disabled context after
MC_CONTEXT_PAUSE_TTL_MS (default 5000; 0 keeps the latch-until-port-recovery
behavior). Reactivation is half-open: the streak resets, so a genuinely
unhealthy RNIC re-trips after another threshold of failures. Fatal async
events clear the breaker and own the inactive state, preventing the TTL
from resurrecting an event-disabled context.
3. Replace the atomic counter with ContextHealthTracker, a mutex/atomic,
clock-injected state machine that tracks the streak, peer set, trip time,
and tripped state together. Tracker callbacks serialize breaker state with
RdmaContext::set_active(), preventing a concurrent trip and recovery reset
from leaving the context inactive with no armed recovery.
Preserve the delayed physical recovery introduced by #2959: PORT_ACTIVE still
starts the 30-second GID recovery probe, clears any breaker trip so its shorter
TTL cannot bypass that delay, and activates the context only after the probe
succeeds.
Add strict parsing for both settings plus hardware-free tests for single- and
multi-peer thresholds, half-open recovery, fatal-event override, transition
races, concurrent access, and configuration boundaries and malformed values.
Signed-off-by: Dao Le <daole@inferact.ai>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* [CI] Rerun flaky UB mock test
* ci: rerun checks
* ci: rerun checks
* fix: simplify RDMA context breaker
* fix: preserve RDMA recovery ownership across GID probes
---------
Signed-off-by: Dao Le <daole@inferact.ai>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* [Store] Fail-stop supervisor on OpLog writer terminal failure (#3923)
When the batch OpLog writer reaches a terminal failure, the HA supervisor now fail-stops serving: mutation admission is closed through admin availability, the RPC delegate and leader label are removed, runtime state becomes standby, and the RPC server is stopped. Terminal notifications are delivered outside writer locks and duplicate notifications are ignored.
Related: Oplog HA PR-F02 design.
Co-authored-by: Yuchen Kou <kouyuchen@approaching.ai>
* [CI/Build] Simplify CI orchestration and share smoke services (#3913)
Simplify CI orchestration and share service lifecycle code across smoke suites. A failed path selector previously skipped the builds but could still leave CI Gate green; the gate now requires selector success and checks each group against its expected success/skipped result.
* [TENT] Add configuration lifecycle foundation (#3836)
* [TENT] Add configuration lifecycle foundation
Refs #3833
* [TENT] Fix configuration lifecycle formatting
* [TENT] Classify HP-TCP configuration lifecycle
* [TENT] Audit lifecycle inventory against source keys
* [TENT] Reject mixed-lifecycle subtree reads
* [TENT] Simplify configuration lifecycle classification
* [TENT] Remove redundant lifecycle from field specs
Store only path and match mode in the runtime allowlist. Return runtime-candidate explicitly for a matching entry and keep the bootstrap-only fallback. Update the inventory test to verify actual classification.
Refs #3833
* [TENT] Reject noncanonical lifecycle view paths
Legacy Config skips empty slash-delimited segments, so raw lifecycle checks can authorize aliases of runtime fields or mixed subtrees. Reject noncanonical paths in the shared view read guard and cover scalar, metrics subtree, and root aliases without changing legacy Config lookup behavior.
Signed-off-by: bugkeep <1921817430@qq.com>
---------
Signed-off-by: bugkeep <1921817430@qq.com>
Co-authored-by: Liqian Lin <jacklin78911@gmail.com>
* [EP] Add NCCL Device API backend for ElasticBuffer (#3449)
* [EP] Add NCCL Device API backend for ElasticBuffer
* [EP] Optimize NCCL rail transport performance
* [EP] Support NCCL 4x4 rail topology
* [EP] Apply clang-format to NCCL backend
* [EP] Make NCCL the default ElasticBuffer backend
* [EP] Remove redundant NCCL template disambiguator
* [EP] Preserve IBGDA hybrid dispatch batching
* [EP] Clarify elastic transport comments
* [EP] Support NCCL 2.30 and 2.31 device layouts
* [EP] Agree on ElasticBuffer transport across ranks
* [EP] Keep NCCL DeviceTransport opt-in
* [TE] Add metadata versioning (#3837)
* [TE] Add metadata versioning
* Fix Copliot issues
* Update local cache when finding latest version
* Fix comments
* [Bugfix][TENT] Serialize lanes sharing a queue pair on its slice queue (#3924)
Co-authored-by: maxlisongsong <maxlisongsong@didiglobal.com>
* [TENT] Fix HP TCP coalescing limits and short READ stalls (#3931)
* fix(tent): bound HP TCP request coalescing by transfer limit
* fix(tent): avoid delayed HP TCP reads and verify rail payloads
* fix(tent): close HP TCP merge and benchmark review gaps
* test(tent): address HP TCP synchronization and CI review
Replace the test-local TSan suppression with the existing worker barrier after WRITE ACKs. Keep the config test's file port occupied by its HTTP server to avoid random-port false failures. Present the existing rail measurements as a table.
* [TENT] Say why the drop and arbitration MLU read different bytes_ahead (#3907)
Co-authored-by: maxlisongsong <maxlisongsong@didiglobal.com>
* [Store] Migrate eviction and concurrency scenarios to the MasterScenario DSL (#3927)
Co-authored-by: Aionw <aione.moe.dev@gmail.com>
Signed-off-by: Rishabh Sinha <rsinha17@terpmail.umd.edu>
* [Store] Split utility code by function (#3587)
* [Store] Route tensor _from writes through multi-buffers (#3722)
---------
Co-authored-by: Teng Ma <stmatengss@gmail.com>
* [Fix] TENT gid lid refresh (#3920)
* [Bugfix][TENT] Refresh RDMA addresses after network changes
Refresh a TENT RDMA context's LID, GID, and GID index on GID_CHANGE,
LID_CHANGE, and port recovery events, then publish the updated device address
to segment metadata before the NIC becomes selectable again.
Keep the context paused during refresh, reject non-active ports, and retire all
endpoints built with the previous address generation before re-enabling the
NIC. Each endpoint pins one immutable local address snapshot so bootstrap,
data QPs, and notification QPs cannot mix generations during a concurrent
refresh.
Add fake-verbs coverage for GID and LID changes, wrong-port events, query
failures, down ports, initially omitted NIC metadata, port recovery, and
endpoint address-generation pinning.
* [Bugfix][TENT] Validate paused state before address refresh
Make RdmaContext::pause report success when it transitions an enabled context
or observes an already-paused context, and reject uninitialized or disabled
states. Gate activateContext address probing and metadata publication on a
confirmed DEVICE_PAUSED state so teardown races cannot publish an address that
can never be resumed.
Add regression coverage for the pause contract and for skipping GID/metadata
refresh when the context cannot enter the paused state.
---------
Co-authored-by: TTThanos <yaozhong.lyz@alibaba-inc.com>
* [TENT] Spin with pause before yielding in TicketLock (#3925)
Co-authored-by: maxlisongsong <maxlisongsong@didiglobal.com>
* [TE] Keep fabric mem Store-only; drop GRC fabric_memory enablement (#3955)
* [TE] Keep fabric mem Store-only; drop GRC fabric_memory enablement
* [TENT] Pin failover policy at transfer submission (#3934)
* [TENT] Pin failover policy at transfer submission
* [TENT] Initialize failover snapshots and share policy defaults
Build the initial snapshot before fallible setup and reject null test publications without replacing the active policy. Share named defaults across snapshots, logical task policies and config parsing, and extend the policy-pinning regression to cover null rejection.
* [TENT] Skip idle GPUs when synchronizing dest CUDA contexts (#3945)
* [TENT] Skip idle GPUs when synchronizing dest CUDA contexts
Do not cudaSetDevice topology cards that have no primary context, and
do not cudaGetDevice when this thread has no current context. Tone
enumerates all visible GPUs without CUDA_VISIBLE_DEVICES; creating
those contexts wastes VRAM on ranks that only use a subset.
Signed-off-by: staryxchen <staryxchen@tencent.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* [TENT] Initialize the CUDA driver once in dest-GPU sync
Call cuInit through std::call_once so topology walks do not re-enter
driver setup. Drop environment-specific comments.
Signed-off-by: staryxchen <staryxchen@tencent.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Signed-off-by: staryxchen <staryxchen@tencent.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* [Store] Validate DummyClient registration of external pinned buffers (#3712)
---------
Co-authored-by: Teng Ma <stmatengss@gmail.com>
* [Bugfix][TransferEngine] Export the requested range for dma-buf registration under CUDA VMM (fixes #2511) (#3538)
* Create a new branch for this commit
Create a new branch for this commit
* Add length parameter to exportDmabuf function
* Update exportDmabuf call to include length parameter
* Update exportDmabuf calls to include buffer size
* Potential fix for pull request finding
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
* Update rdma_context.cpp
style: satisfy clang-format (collapse single-statement if)
---------
Co-authored-by: Feng Ren <alogfans@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
* [Py] Fix batch_id leak on timeout and submit failure paths (#2350)
* [Integration] Fail Python init when allocator support is unavailable (#2404)
* fix: fail Python allocator init safely
* Mark initMemoryAllocator static; it is only used inside this TU
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
---------
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
* TENT: deliver self-targeted notifications in-process (#3464)
* TENT: add the in-process queue for self-targeted notifications
Hold notifications whose target is LOCAL_SEGMENT_ID in the engine itself. The queue is drained by receiveNotification() in the next commit; no behavior change on its own.
* TENT: deliver self-targeted notifications in-process
The data plane short-circuits transfers whose target is LOCAL_SEGMENT_ID to the local path, but notifications had no such branch: sendNotification() handed every target to the first transport that supports notifications. Over RDMA that reaches the local pseudo-endpoint, which never performs the bootstrap handshake and so never owns a notification QP - the send fails ("Notification QP not connected", rdma_transport.cpp -> endpoint.cpp) and a receiver polling for the notification waits forever.
Observed end to end through the NIXL Mooncake plugin: the first intra-agent transfer carrying a notification hangs its receiver until timeout. The combination went unnoticed because the intra-agent DRAM unit test upstream is commented out and the VRAM variant only compiles with CUDA on a GPU machine.
Deliver self-targeted notifications in-process instead: queue them in the engine on send, and drain that queue in receiveNotification() after polling the transport. This mirrors the data plane's local short-circuit, works for every transport combination, and needs no wire at all - including builds where no transport supports notifications, where a queued self-notification is now delivered rather than rejected.
* TENT: cover self-targeted notification delivery with a unit test
local_notification_test pins the contract without needing a NIC or a metadata server: in-order exactly-once delivery of self-targeted notifications, resolution of the engine's own advertised name (in p2p mode the normalized "ip:port", not the configured local_segment_name) and of the empty name to LOCAL_SEGMENT_ID, and delivery with every transport disabled.
* TENT: address review on self-targeted notification delivery
- receiveNotification() clears the output list on entry. Callers reuse one
vector across polls, so appending to leftovers redelivered whatever the
caller had already handled.
- A poll that delivered something reports success. Notifications appended to
the list have been consumed from their queue, so the caller has to see
them; a persistent transport error still surfaces on the next poll, which
comes back empty. The local queue is drained regardless of the transport
poll's outcome, since in-process delivery does not depend on it.
- local_notification_test gains the case the fix is actually about: an
intra-agent transfer carrying a notification, i.e. submitTransfer(...,
notifi) -> maybeFireSubmitHooks() -> sendNotification(LOCAL_SEGMENT_ID).
It routes notifications through a stand-in transport that advertises
support and fails to send, which is what the RDMA local pseudo-endpoint
does. Without that stand-in a real transport answers on loopback and the
notification round-trips even without the fix, so the test would not pin
the regression; verified to fail with the local short-circuit removed and
pass with it.
- makeConfig disables the remaining transports (sunrise_link, mpcomm, tpu,
ub) so the no-notification-transport case is genuinely transport-free, and
both delivery tests now poll twice through the same vector.
tent suite: 37/37 passed.
* TENT: notify each submit-hook target exactly once
A submit hook is only marked fired once every target has taken delivery,
and it reruns from every later poll while any of them still fails. The
loop restarted from the first target each time, so a target that had
already been notified was notified again on each of those polls.
Self-targeted delivery is what makes that reachable. Before it, the local
leg failed alongside the remote one and nothing was delivered at all; now
it always succeeds, so a batch that notifies both the local segment and a
peer that stays unreachable re-queues the local copy on every poll -
unbounded growth in the local queue, and one duplicate "transfer
complete" per poll for the receiver. Drop each target from the hook as it
takes delivery instead, so a rerun only retries the ones still failing.
local_notification_test covers it with two loopback engines: a two-target
notified transfer whose peer leg cannot be notified, polled while it
fails and again after it recovers. The notification must arrive exactly
once across both phases; without this change it arrives seven times.
Also from the same review pass:
- Document that receiveNotification() clears the caller's vector on
entry. That was introduced by the previous commit and is the opposite
of the classic engine's getNotifies(), which appends, so it is worth
stating rather than leaving for a caller to discover.
- Pin enable_progress_worker and enable_runtime_queue off in the test
config. The progress worker fires submit hooks from its own thread,
which races the stand-in transport's counter that these tests read from
the main thread, and the runtime queue turns it on implicitly.
tent suite: 32/32 passed.
* TENT: attempt every submit-hook target on every pass
Dropping a target as it took delivery still stopped the loop at the first
send failure. hook.targets is an unordered_set, so every target behind the
failing one went unattempted, and because the order is stable across polls
it stayed that way. A self-targeted notification cannot fail, so an
unrelated peer being down would hold one back and block its receiver until
that peer recovered - the liveness of in-process delivery depended on
iteration order.
Continue past failures instead: erase each target that takes delivery, keep
the rest for the next pass, and mark the hook fired once none are left.
SelfNotificationIsNotHeldBackByAnUnreachablePeer covers it with two
unreachable peers and a local target. The delivery assertion alone would
only fail when the order happens to put a peer first, so the test also
counts send attempts: every pass must attempt both failing peers. Against a
version that stops at the first failure that count is 4 over 4 passes
instead of 8, which does not depend on the order.
* Re-trigger CI: tent-ci (ub-mock) hit a known race in RuntimeQueueDispatch.ProgressWorkerRefillsWindowFromTransportNotify
---------
Co-authored-by: Frank <xiaodouzi6661@gmail.com>
* [Store] Introduce SegmentPool catalog and lifecycle (#3705)
Part of #3363; architectural work tracked by #3360. The placement foundation from #3704 is already in main. This PR adds the next layer: a SegmentPool that coordinates mounted-region metadata, allocator resources, placement eligibility, and lifecycle transactions under one lock. It also simplifies the placement interfaces introduced by that foundation.
For example, preparing an unmount removes a region from allocation immediately while retaining the driver resource for cleanup. An immediate transaction can commit or roll back; a graceful transaction keeps existing replicas readable until finalization. Catalog state, driver activation, and placement membership change together under the Pool write lock.
* [Store] Structure NvmeKvConnectorConfig environment settings (#3948)
Signed-off-by: Schatten <czhengt@qq.com>
* [Store] Structure NoFRegisterConfig environment settings (#3949)
Co-authored-by: Aoi <aione.moe.dev@gmail.com>
Signed-off-by: Schatten <czhengt@qq.com>
* [Store] Structure MmapArenaConfig environment settings (#3956)
Signed-off-by: Schatten <czhengt@qq.com>
* [Store] Structure ClientNumaConfig environment settings (#3966)
Signed-off-by: Schatten <czhengt@qq.com>
* [Store] Structure CxlSegmentConfig environment settings (#3967)
Co-authored-by: Aoi <aione.moe.dev@gmail.com>
Signed-off-by: Schatten <czhengt@qq.com>
* [Store] Wire opt-in batch OpLog snapshots into standby runtime (#3950)
Co-authored-by: Yuchen Kou <kouyuchen@approaching.ai>
* [Store] Add recoverable client liveness and asynchronous offboarding (#2991)
* feat(store): add client liveness state and compatibility config
* feat(store): gate client-owned resources by liveness
* feat(store): retry terminal client offboarding
* feat(store): rebuild liveness across HA restore
* refactor(store): remove client liveness leftovers
* refactor(store): separate allocator registration lifecycle
* fix(store): preserve multi-segment client resources
* fix(store): preserve client liveness invariants
* test(store): align liveness and OpLog scenarios
---------
Co-authored-by: zhuyuening <zhuyn25@mails.tsinghua.edu.cn>
* [Store] Sweep stale handles in bounded write-lock batches (#3593)
Signed-off-by: pjdurden <prajjwalchittori1@gmail.com>
* [TransferEngine] Add POSIX SHM transport for same-host DRAM copies (#3918)
Classic TE still hairpinned same-host DRAM through RDMA/TCP loopback.
Add a first-class ShmTransport (HIP-style) so shm-backed buffers can
memcpy locally, with MC_FORCE_SHM=1 and optional RDMA/TCP coexistence
when ENABLE_MULTI_PROTOCOL is on. Pin relocate-cache mappings across
memcpy so prune/cap cannot munmap an in-flight copy.
RFC: https://github.com/kvcache-ai/Mooncake/issues/3912
Signed-off-by: staryxchen <staryxchen@tencent.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* [TransferEngine] Prefer THP for shared-segment mmap without a HugeTLB pool (#3952)
POSIX shm mapped first and only hinted THP afterward, so large HostRegister
MRs often stayed on 4KiB pages. Unnamed memfd plus a 2MiB-aligned VA and
MADV_HUGEPAGE before fault lets TP ranks share the same THP-backed pages.
Co-authored-by: youyx <default>
Co-authored-by: Cursor <cursoragent@cursor.com>
* [EP] Support NCCL ElasticBuffer recovery after rank replacement (#3963)
* feat(ep): rebuild NCCL state after rank recovery
* fix(ep): allow reserved PG capacity during NCCL recovery
Validate the buffer's logical ranks independently of reserved PG mask capacity. Exchange expected world sizes before membership collectives so actual resizing is rejected consistently, and add unit and GPU recovery regressions.
* fix(ep): address NCCL recovery review feedback
Recognize the mooncake-cpu backend and use CPU metadata tensors for NCCL bootstrap. Call the always-registered NCCL capability binding directly. Move recovery integration tests into the EP suite while reusing the PG worker harness, add CPU-PG recovery coverage, and document the test command.
* [TENT] Export EGM host memory over multi-node NVLink (opt-in) (#3978)
On Grace-Blackwell NVLink domains GPUs can address peer nodes' host DRAM
through Extended GPU Memory: host memory created with
cuMemCreate(CU_MEM_LOCATION_TYPE_HOST_NUMA) exports as a fabric handle
like device memory. The mnnvl transport only exported device buffers, so
CPU-resident data always crossed the NIC.
Add transports/mnnvl/egm (env MC_MNNVL_EGM, default off). When enabled:
- allocateLocalMemory("cpu:<numa>") returns an EGM allocation mapped for
every local GPU and the CPU, and the engine's default host allocation
goes through the mnnvl transport;
- addMemoryBuffer exports host buffers that are EGM allocations, other
host memory keeps the cudaHostRegister path;
- relocateSharedMemoryAddress imports any exported peer buffer, not only
cuda:* locations;
- the transport advertises dram_to_dram and gpu_to_dram.
The engine's mnnvl staging policy keys on whether the peer actually
exported the host buffer, so peers registering plain host memory keep
the staging path they had before this flag existed.
The copy stream for host buffers now runs on a GPU attached to the
buffer's NUMA node instead of parsing "cpu:N" as CUDA device N.
Handle lifetime: addMemoryBuffer drops the reference taken by
cuMemRetainAllocationHandle once the buffer is exported (host and device
paths), and transport-owned VMM allocations release the creator handle
after mapping and remember the granularity-rounded extent, so
unregister + freeLocalMemory actually return the memory.
Claude-Session: https://claude.ai/code/session_01USLEbvBiLAhKWPQrvVAMC9
Signed-off-by: JinYan Su <751080330@qq.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* [Store] Release GIL during client setup (#3947)
Release the Python GIL while setup_real and setup_internal perform native initialization so other Python threads can continue running during long memory registration.
Signed-off-by: leichao.lc <leichao.lc@antgroup.com>
Co-authored-by: leichao.lc <leichao.lc@antgroup.com>
* [Reshard] Add KV cache reshard planning (#3564)
* feat: Add kv reshard
* fix(reshard): repair KV merge integration
* fix
* Phase1
* put lowering to mooncake
* move lowering
* feat(reshard): add runtime-to-runtime KV cache transfer execution
- Harden snapshot, placement, and logical plan validation
- Add operation-scoped resolved bindings and bounded transfer lowering
- Add TE batch execution and writer/target completion contracts
- Preserve existing SGLang PD planning interfaces
- Add contract and native CPU/GPU TCP/RDMA transfer tests
* fix(reshard): tighten KV cache snapshot and memory validation
- Reject mismatched binding snapshots while preserving legacy PD calls
- Validate derived transfer byte sizes against uint64 bounds
- Scope address checks by runtime instance, worker, and device
- Reject conflicting backing allocations while allowing shared storage
- Consolidate boundary tests and remove redundant assertions
Signed-off-by: huanglong <huanglong@linux.alibaba.com>
* fix(reshard): quarantine KV executor on registration failures
Signed-off-by: huanglong <huanglong@linux.alibaba.com>
---------
Signed-off-by: huanglong <huanglong@linux.alibaba.com>
* [CI/Build] Add self-hosted MI350X ROCm external PD tier (#3338)
Co-authored-by: Aionw <aione.moe.dev@gmail.com>
* [Bugfix] Preserve calling thread CPU affinity across SPDK env init (#3904)
* Preserve calling thread CPU affinity across SPDK env init
rte_eal_init registers the calling thread as DPDK's main lcore and
restricts its CPU affinity to the EAL core mask (default "0x1"). Linux
threads inherit the creator's affinity mask, so without a restore every
thread spawned after setup (e.g. sglang's prefetch/backup workers) would
be pinned to a single core, sharply degrading KV read/write throughput.
Save and restore the caller's affinity around spdk_env_init so default
behavior is unchanged.
Signed-off-by: tan changzhi <544463199@qq.com>
Co-Authored-By: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
* [Store] Gate affinity save/restore to Linux in SPDK env init
pthread_getaffinity_np/pthread_setaffinity_np and cpu_set_t are
glibc-specific; without the guards the affinity fix would break
non-Linux USE_NOF builds. Linux behavior is unchanged; other
platforms keep the pre-fix behavior.
Co-Authored-By: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
---------
Co-authored-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
* [Store] Add DFS read/write metrics and diagnostic logging (#3855)
Co-authored-by: sunbaoshan <sunbaoshan@baidu.com>
Co-authored-by: root <root@gajl-bbc-onlinec-com-1398657.gajl.baidu.com>
* [Store] Support online DFS shard capacity expansion (#3984)
* [Store] Support online DFS shard capacity expansion
* [Store] Isolate lazy DFS shard initialization
* [P2PStore] Fix GetReplica metadata recheck retries (#3921)
* [Store] Fail KV benchmarks when a worker raises (#3938)
Propagate lane failures after joining all workers so the CLI cannot report
success for an incomplete phase or close buffers still used by another lane.
Signed-off-by: Zhe Li <2843409461@qq.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
* [TE] Report buffered JSONL write failures in tebench (#3992)
Co-authored-by: jipengtian <jipengtian@xiaohongshu.com>
* [Bugfix][TransferEngine] Invalidate stale remote segments in syncSegmentCache (#3072)
* [Bugfix][TransferEngine] Invalidate stale remote segments in syncSegmentCache
syncSegmentCache logged "segment X is now invalid" when a cached remote
segment's backend fetch failed but never erased the stale descriptor, so
peers kept routing transfers to dead/unmounted peers until the caching node
restarted. The existing erase sites were local-only paths and did not cover a
remote segment that vanished from the backend.
Track a per-segment consecutive fetch-failure streak (guarded by
segment_lock_) in the apply phase. After kStaleSegmentFailureThreshold (=2)
consecutive failures, erase the entry from both segment_name_to_id_map_ and
segment_id_to_desc_map_, mirroring removeSegmentDesc. A successful fetch
resets the streak, so a single transient backend blip (curl timeout / etcd
hiccup, which the plugin get() collapses with "not found" into false) does
not purge the cache and strand in-flight transfers.
Adds a hardware-free unit test (fake MetadataStoragePlugin) that is
red-on-master (stale entry never invalidated) and green-on-branch.
* style: clang-format-20 the stale remote-segment invalidation test
* [TransferEngine] Distinguish NOT_FOUND from UNAVAILABLE in metadata plugin
MetadataStoragePlugin::get() collapsed a missing key with a transient
backend error into a single false, so two consecutive backend timeouts
looked identical to an authoritative removal and could purge the entire
live cache during a metadata-service outage.
Add GetResult enum (kFound/kNotFound/kUnavailable) and getWithStatus()
virtual with a default that delegates to get() (preserving behavior for
plugins that don't distinguish). Override in RedisStoragePlugin (nil
reply -> kNotFound, command error -> kUnavailable) and HTTPStoragePlugin
(HTTP 404 -> kNotFound, curl/5xx -> kUnavailable).
syncSegmentCache now advances the invalidation streak only for kNotFound;
transient failures (kUnavailable) leave the cached entry untouched and
are retried next cycle.
Add BackendOutageDoesNotInvalidateCache regression: 5 consecutive
transient failures must not invalidate the cached entry.
* fix(format): align getSegmentDesc declaration with clang-format-20 style
clang-format-20 fills the first line up to the 80-column limit instead
of breaking immediately after the opening paren.
Signed-off-by: supermario_leo <leo.stack@outlook.com>
* [TransferEngine] Fail the default getWithStatus() mapping closed
The default getWithStatus() mapped every get() failure to kNotFound, so
a plugin that cannot distinguish absence from an outage (both etcd
implementations do not override it) would let syncSegmentCache() evict
a still-live cache entry during an etcd outage -- reintroducing the
exact availability bug this PR fixes.
Flip the default to fail closed: an ambiguous false is kUnavailable,
and only plugins that can prove an authoritative absence may report
kNotFound. Both etcd plugins now override getWithStatus(): the legacy
SyncClient treats an OK range response with an empty payload as
kNotFound, and the wrapper build treats ret == 0 with no payload as
kNotFound; RPC errors, timeouts and malformed payloads are
kUnavailable.
Add DefaultPluginMappingFailsClosed regression exercising the inherited
default path (a plugin without an override), which stays red until the
fail-closed mapping lands.
Signed-off-by: supermario_leo <leo.stack@outlook.com>
---------
Signed-off-by: supermario_leo <leo.stack@outlook.com>
* [CI/Build] Fix build-musa: map cudaDevAttrMultiProcessorCount (#3999)
* [Store] Complete N08 post-launch snapshot validation (#3991)
Co-authored-by: Yuchen Kou <kouyuchen@approaching.ai>
* [Bugfix][Store] Cover HotStandbyService behaviors with deterministic fakes (#3687)
Co-authored-by: YDMY_Kirito <hantenglong@kylinos.cn>
Co-authored-by: Cursor <cursoragent@cursor.com>
* [TENT] Rotate HP TCP lanes independently per peer (#3987)
* [Store] Avoid blocking HA reattachment on master connection retries (#3743)
* [P2P-Store] Back off transfer status polling after a spin budget (#3851)
Signed-off-by: chlins <chlins.zhang@gmail.com>
* [TENT] Retire idle HP TCP client connections (#3988)
* [Conductor] conductor: prefix index types and hash strategy (#3879)
Co-authored-by: Misak2333 <167268798+Misak2333@users.noreply.github.com>
* [Store] Optimize CRC-64 checksum calculation with slice-by-8 (#3922)
* [Bugfix][Store] Check read leases after grouped batch gets (#3996)
* [Docs] Clarify Store parallel tensor API selection and limits (#3959)
* [Docs] Clarify Store parallel tensor API selection and limits
* [Docs] Fix parallel tensor API table format and spell check
Replace torch.arange with torch.linspace so typos CI accepts the example,
use a standard padded pipe table, and switch the planner cross-link to a
Sphinx {ref} target.
Co-authored-by: Teng Ma <stmatengss@users.noreply.github.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Teng Ma <stmatengss@users.noreply.github.com>
* [Store] Fan verified bytes out to duplicate keys in batch gets (#3929)
The destination slice maps are keyed by string, so a repeated key
collapses: both indices transfer into the last op's buffer, the
per-index checksum validates against that one filled buffer, and the
first op comes back SUCCESS with its own never-written (recycled,
stale) memory (#3876). The same collapse exists in the
batch_get_buffer / batch_get_into / batch_get_into_multi_buffers
memory paths, the DISK/DFS temp-buffer paths, and both LOCAL_DISK
offload paths.
Transfer each unique key once (first occurrence as primary), then fan
the verified bytes out to each duplicate's own destination locally.
The dedup plan and the result propagation now live in one shared
layer (batch_read_fanout.h: PlanDuplicateKeys, FanOutDuplicates), so
every read path behaves identically: duplicates of a failed primary
inherit the primary's error instead of a pre-set success, and a size
mismatch between same-key destinations fails the duplicate loudly
everywhere.
The fan-out copy is device-aware end to end: RuntimeAccelerator gains
CopyAuto (host/device on either side, memcpy fallback) and
CopySlicesMaybeDevice walks scatter-gather lists whose chunking may
differ, so MEMORY/LOCAL_DISK duplicates with GPU-resident caller
buffers no longer go through a plain memcpy. New unit tests pin the
plan, the slice walker (including mismatched chunking and
underflow), the fan-out contract, and CopyAuto's direction choice;
the existing integration test still covers the end-to-end behavior.
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
* [Bugfix][TE] Fix Ascend LocalCopyEngine stream/context for multi-device async copy (#4026)
* [Bugfix][TE] Fix Ascend LocalCopyEngine stream/context for multi-device async copy
* docs: update README AgentX news (#4011)
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: stmatengss <11641725+stmatengss@users.noreply.github.com>
* [TENT] fix: TCP transport implicitly creates CUDA contexts (#3973)
* [TENT] fix: TCP transport implicitly creates CUDA contexts
* Re-trigger CI (build-wheel-rocm 3.12 flake)
* [TransferEngine] Enable MACA IBGDA using the shared mlx5 transport (#3997)
* [Store] Align BucketBackendConfig with config layout (#3957)
Signed-off-by: Schatten <czhengt@qq.com>
* [Store] Structure NvmeKvIoConcurrencyConfig environment settings (#4012)
Signed-off-by: Schatten <czhengt@qq.com>
* [Store] Structure HugepageConfig environment settings (#3983)
Signed-off-by: Schatten <czhengt@qq.com>
* [Store] Structure ClientAutoDiscoveryConfig environment settings (#3900)
Co-authored-by: Aoi <aione.moe.dev@gmail.com>
Signed-off-by: Schatten <czhengt@qq.com>
* [Store] Structure SpdkControllerConfig environment settings (#3982)
Co-authored-by: Aoi <aione.moe.dev@gmail.com>
Signed-off-by: Schatten <czhengt@qq.com>
* [Store] Structure RpcProtocolConfig environment settings (#3976)
Co-authored-by: Aoi <aione.moe.dev@gmail.com>
Signed-off-by: Schatten <czhengt@qq.com>
* [Doc] Backport vLLM-style notes for large changes (#4032)
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: stmatengss <11641725+stmatengss@users.noreply.github.com>
* [Conductor] Add PR classification and automatic labeling (#4010)
* [TENT] Revoke registrations removed by failed batch rollback (#4002)
* [TENT] Revoke registrations removed by failed batch rollback
Hand rollback-to-zero descriptors from SegmentTracker to the engine after
metadata publication, then finish transport cleanup before metadata sync
and return. Preserve the original registration error and normal unregister
behavior.
Add tracker and real-engine regressions for cleanup ownership, failures,
metrics, lease draining, replacement survival and stale HP TCP writes.
* [TENT] Remove unused rollback test callback
* [TENT] Remove added rollback regression fixtures
* [TENT] Reuse callback handoff for failed registration rollback
* [TENT] Fix standard TCP completion, permissions and task accounting (#4001)
* [TENT] fix TCP task lifetime after terminal publication
* [TENT] fail TCP tasks when worker operations throw
* [TENT] enforce registered memory permissions in TCP RPC handlers
* [TENT] preserve original byte counts for merged public tasks
* [TENT] evaluate batch notifications per submit interval
* [TENT] bound standard TCP RPC payloads and request merging
* [TENT] keep TCP merge route preflight free of invalidation
* [TENT] add focused TCP correctness regressions
* [Store] Support RealClient CUDA tensor reads (#3723)
---------
Co-authored-by: Teng Ma <stmatengss@gmail.com>
* [TENT] Refresh TCP endpoints without replaying uncertain writes (#4006)
* [TENT] Refresh TCP endpoints without replaying uncertain writes
* [TENT] Remove added TCP endpoint refresh regression tests
* [TENT] Clarify TCP transfer failure log
* [TENT] Refresh HP TCP endpoints after pre-send connection failures (#4007)
* [TENT] Refresh HP TCP endpoints after pre-send failures
* [TENT] Remove added HP TCP endpoint refresh regression test
* [TENT] Limit HP TCP pre-send refresh to endpoint errors
* TENT: derive notification QP attributes from endpoint params (#4064)
The notify QP hard-coded the verbs tutorial values for its path and retry
attributes while the data QPs of the same endpoint take them from
EndPointParams (setupOneQP), so the knobs under transports/rdma/endpoint
never reached the notify path. Read them from the same params, so both
QPs of an endpoint run on one budget and one configuration:
attribute before after (default)
path_mtu 4096 path_mtu (4096, clamped to the port MTU)
min_rnr_timer 0x12 min_rnr_timer (12: 0.64 ms, was 5.12 ms)
timeout 0x12 send_timeout (14: 67 ms/attempt, was 1.07 s)
retry_cnt 7 send_retry_count (7)
rnr_retry 7 7, pinned (see below)
max_*_rd_atomic 1 1 (the notify QP only carries SEND/RECV)
The EndPointParams defaults themselves are unchanged. Nominally this
brings the time to report a dead peer path on the notify QP from
4.096us * 2^18 * 8 ~= 8.6 s down to the ~0.54 s budget the data QPs
already have; on a BlueField-3 integrated ConnectX-7 (VF) the same
change measured 14.9 s -> 3.6 s, against 3.7 s for the data QPs with the
same parameters.
path_mtu following params is a side effect worth having: the notify QP
hard-coded IBV_MTU_4096 and drifted from the data QPs on a fabric that
configures a smaller transports/rdma/endpoint/path_mtu. min_rnr_timer
moving from 0x12 (5.12 ms) to the default 12 (0.64 ms) shortens the RNR
retry interval: a burst of 2000 standalone notifications against the 256
posted recv slots went from p50 15.3 ms to 1.6 ms. rnr_retry stays
pinned at 7 (unbounded) and does not read send_rnr_count: RNR NAKs are
ordinary flow control on the notify QP, the endpoint's only SEND/RECV
QP, and must never retire it together with its data QPs. The
address-vector fields (hop_limit, flow_label, src_path_bits, PSNs) are
left as they were, and the path-class completion handling from #3475 is
untouched.
Co-authored-by: xiaodouzi666 <xiaodouzi6661@gmail.com>
* [CI/Build] Enable KV events in wheels with ZeroMQ dependencies (#4060)
---------
Signed-off-by: Ishan Dhanani <ishandhanani@gmail.com>
Co-authored-by: Ishan Dhanani <ishandhanani@gmail.com>
* [Bugfix] Fix GIL misuse in TENT Python guard exits (#4051)
* [Bugfix] Fix GIL misuse in TENT Python guard exits
* [TENT] Cover normal and exceptional Python guard exits
* [Store] Support DeepSeek-V4.1-Flash Engram with local and Store-backed byte tables (#4065)
* [Store] Support multi-layer byte Engram tables with caller-owned output
Replace the float32-specific Engram layout with a required positive row_bytes field and retain fixed per-layer, per-head keys. One EngramStore manages a map of layer layouts over a shared Store client; reads, writes and removal select a layer explicitly.
Expose lookup_into as the only read interface, writing directly into caller-owned, pre-registered uint8 output without allocating or registering output internally. Pass replication configuration through populate, validate row sizes and buffer layouts, and update integration tests, the benchmark, and API documentation.
Validation: Store binding build, 17 integration tests, small benchmark smoke, scoped pre-commit and documentation HTML build passed.
---------
Signed-off-by: Teng Ma <stmatengss@gmail.com>
* [CI/Build] Retry transient PUT failures in SSD eviction test (#3875)
* [Bugfix][TENT] Plan RDMA slices from the block size: no empty slices, and fold short tails (#3981)
* [Bugfix][TENT] Take the slice count from the block size so no slice is empty
* [Bugfix][TENT] Fold a short slice tail into the slice before it, not into the block size
* [Bugfix][TENT] Charge the folded slice tail so allocate and release balance
---------
Co-authored-by: maxlisongsong <maxlisongsong@didiglobal.com>
* [TENT] Drain standard TCP workers during quiesce (#4005)
* [TENT] Drain standard TCP workers before batch teardown
* [TENT] Remove added TCP teardown regression tests
* [TENT] Synchronize TCP admission with quiesce
* [Store] Add a per-tenant eviction watermark so quota pressure evicts instead of rejecting (#3805)
* [Store] Structure RedisConnectionConfig environment settings (#4036)
Signed-off-by: Schatten <czhengt@qq.com>
* [TENT] Add CUDA coverage for TCP data-path round trips (#4082)
* [Bugfix][TENT] Keep engines alive for Python allocation guards (#4087)
* [TransferEngine] Enrich Ascend init and register-mem error logs with context (#4084)
* [TransferEngine] Enrich Ascend init and register-mem error logs with context
Include endpoint index/counts and register addr/length so engine mismatches and ADXL register failures are easier to diagnose.
Co-authored-by: Cursor <cursoragent@cursor.com>
* [TransferEngine] Apply clang-format to Ascend error log lines
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* perf(store): avoid redundant query result copies (#4085)
Signed-off-by: waizuichougou <2082431897@qq.com>
* [TENT] Add Intel XPU platform to the Transfer Engine (#4030)
Add an XpuPlatform to the TENT runtime so Mooncake can stage transfers
through Intel XPU (GPU) device memory. XpuPlatform extends CpuPlatform:
host DRAM allocation, NUMA probing, and host<->host copies are inherited
unchanged, and only the XPU-device-aware operations are overridden.
Implementation is a native, direct-link SYCL build: the platform sources
include <sycl/sycl.hpp> and link libsycl directly via a private
XpuSyclBackend (device enumeration with GPU-preferred/any-device
fallback, sycl::malloc_device USM allocation, an interior-pointer
registry so chunked staging addresses classify correctly, and
VRAM<->host memcpy). Because SYCL requires the Intel DPC++ compiler,
USE_XPU=ON now requires USE_TENT=ON and fails fast unless the build is
configured with icpx (CXX=icpx); -fsycl is applied to platform_xpu and
propagated to everything that links it.
- runtime: MTYPE_XPU memory type + USE_XPU loader branch
- platform: XpuPlatform (probe/allocate/free/copy/getMemoryType/
getLocation) backed by XpuSyclBackend
- build: USE_XPU option/gate in common.cmake; platform_xpu with -fsycl
- test: tent_xpu_platform_test drives the real backend (alloc -> H2D ->
D2H -> byte-equality -> free, host/interior pointer classification),
skipping when no SYCL device is present; runs against the OpenCL CPU
device on GPU-less hosts
- docker/xpu.Dockerfile: build image based on intel/pytorch:xpu with the
oneAPI DPC++ compiler added, so it both builds and runs the platform
Test on Intel B60 platform.
* [Bugfix][TENT] Fix Python send_probe under MC_USE_TENT (#4081)
* [Bugfix][TENT]…
Description
Add a trusted self-hosted ROCm external prefill/decode CI tier for the two-node MI350X cluster.
This revision replaces the previously staged T-One controller because T-One does not provide AMD GPUs. The change:
workflow_dispatch, a push tomain, or the maintainer-controlledrun-e2e-cipull_request_targetpath;ionic_0-ionic_3, selects GID index 1, and forces HCA transport so a missing RDMA path cannot silently fall back to TCP;The host profile is documented in
scripts/tone_tests/rocm_runner.env.example. No T-One credential or long-lived GitHub token is required by the ROCm job.The organization runner
mooncake-rocm-mi350x-controlleris registered with labelsamd,rocm,gfx950, andmooncake-pd, but its service is intentionally disabled. Before activation, an organization administrator must place it in a runner group restricted tokvcache-ai/Mooncakeand the trusted CI/E2E workflows. This PR remains a draft until a trusted end-to-end run and burn-in complete.Module
mooncake-transfer-engine)mooncake-store)mooncake-ep)mooncake-pg)mooncake-integration)mooncake-p2p-store)mooncake-wheel)mooncake-common)Type of Change
How Has This Been Tested?
Static validation:
The targeted Actionlint, YAML parser, shell syntax, codespell, pre-commit, local-wheel input, and whitespace checks pass. The skipped pre-commit hooks would report unrelated pre-existing text or normalize legacy shell files outside this change.
Host validation:
mooncake-ciuser;ibv_devinfoopensionic_0throughionic_3inside the least-privilege controller container; andThe two-node Mooncake/SGLang/vLLM end-to-end Actions job remains pending the trusted upstream activation described above.
Test results:
Checklist
./scripts/code_format.sh(not applicable; no C/C++ files changed)pre-commit run --all-filesand all hooks pass (targeted relevant hooks pass)AI Assistance Disclosure
OpenAI Codex helped analyze the CUDA/ROCm CI gap, adapt the integration scripts and workflow for a self-hosted MI350X controller/worker pair, validate resource isolation, and prepare this draft. The human submitter will review and defend every changed line before the PR is marked ready.